Singular value decompositions (SVD)

Lecture 31

Author
Affiliation

Minjae Park

Auburn University
MATH 2660 - Spring 2026

Published

April 1, 2026

Recap

Orthogonal matrices

  • Let \(\{\vec{u}_1, \ldots, \vec{u}_n\}\) be an orthonormal basis of \(\mathbb{R}^n\), meaning
    • \(\vec{u}_i \cdot \vec{u}_j = 1\) if \(i=j\)
    • \(\vec{u}_i \cdot \vec{u}_j = 0\) if \(i \neq j\)
  • Let \(P = [\vec{u}_1 \ \cdots \ \vec{u}_n]\) be the matrix whose columns are these vectors.
  • Then \[P^T P = P P^T = I_n \quad \text{and hence} \quad P^{-1} = P^T\]
  • Such a matrix is called an orthogonal matrix.

Symmetric matrices

  • Let \(A\) be an \(n \times n\) matrix. If \(A^T = A\), then \(A\) is symmetric.
  • Spectral Theorem (key fact):
    • A symmetric matrix is diagonalizable.
    • All eigenvalues are real.
    • There exists an orthonormal eigenbasis.
  • Suppose \(\{\vec{u}_1, \ldots, \vec{u}_n\}\) is an orthonormal eigenbasis with \[A \vec{u}_i = \lambda_i \vec{u}_i\]
  • Let \(P = [\vec{u}_1 \ \cdots \ \vec{u}_n]\) and \(D = \mathrm{diag}(\lambda_1, \ldots, \lambda_n)\).
  • Then (which is called an orthogonal diagonalization) \[A = P D P^{-1} = P D P^T\]

Singular value decomposition

Motivation

  • If an \(n \times n\) matrix \(A\) is diagonalizable, we understand its action easily:
    • it stretches space along eigenvector directions.
  • But:
    • Not every matrix is diagonalizable as we have seen in several examples.
    • Non-square matrices (\(n \times m\)) don’t even have eigenvalues in the usual sense.
  • Question: can we still describe the geometric action of any matrix?
  • Answer: Singular Value Decomposition (SVD).

Eigenvalues of \(A^T A\)

  • Let \(A\) be an \(n \times m\) matrix, not necessarily square.
  • Consider the matrix \(A^T A\).
  • Then:
    • \(A^T A\) is an \(m \times m\) matrix,
    • \(A^T A\) is symmetric, since \[ (A^T A)^T = A^T (A^T)^T = A^T A \]
    • therefore \(A^T A\) is diagonalizable with real eigenvalues.
  • In fact, something even better is true:
    • every eigenvalue of \(A^T A\) is nonnegative.
  • Key idea: even if \(A\) itself is not square or not diagonalizable, the matrix \(A^T A\) is always well behaved.

Why are the eigenvalues nonnegative? (optional)

  • Let \(\lambda\) be an eigenvalue of \(A^T A\), and let \(\vec{v} \neq \vec{0}\) be a corresponding eigenvector.
  • Then \[ A^T A \vec{v} = \lambda \vec{v} \]
  • Multiply both sides on the left by \(\vec{v}^T\): \[ \vec{v}^T A^T A \vec{v} = \lambda \vec{v}^T \vec{v} \]
  • The right-hand side is \[ \lambda \vec{v}^T \vec{v} = \lambda \left\lVert \vec{v} \right\rVert^2 \]
  • For the left-hand side, regroup as \[ \vec{v}^T A^T A \vec{v} = (A \vec{v})^T (A \vec{v}) = \left\lVert A \vec{v} \right\rVert^2 \]
  • Therefore \[ \lambda \left\lVert \vec{v} \right\rVert^2 = \left\lVert A \vec{v} \right\rVert^2 \ge 0 \]
  • Since \(\vec{v} \neq \vec{0}\), we have \(\left\lVert \vec{v} \right\rVert^2 > 0\).
  • Hence \[ \lambda \ge 0 \]

Singular values

  • Let \(A\) be an \(n \times m\) matrix.

  • Let \(\lambda_1, \ldots, \lambda_m\) be the eigenvalues of \(A^T A\).

  • Since these eigenvalues are all nonnegative, we can take their square roots.

  • The singular values of \(A\) are defined by \[ \sigma_i = \sqrt{\lambda_i} \ge 0 \]

  • By convention, we order them in descending order: \[ \sigma_1 \ge \sigma_2 \ge \cdots \ge \sigma_r \ge \sigma_{r+1} = \cdots = \sigma_m = 0 \]

  • Important fact: the number of nonzero singular values equals the rank of \(A\), denoted \(r = \mathop{\mathrm{rank}}(A)\).

Example

  • If \[ A = \begin{bmatrix} 1 & 0\\ 0 & 2\\ 0 & 0 \end{bmatrix} \] then \[ A^T A = \begin{bmatrix} 1 & 0 & 0\\ 0 & 4 & 0\\ 0 & 0 & 0 \end{bmatrix} \]
  • Its eigenvalues are \(1, 4, 0\).
  • Taking square roots gives singular values: \[ 1,\ 2,\ 0 \]
  • Reordering in descending order: \[ \sigma_1=2,\quad \sigma_2=1,\quad \sigma_3=0 \]
  • Since there are two nonzero singular values, \(\operatorname{rank}(A)=2\).

Singular value decomposition (SVD)

Let \(A\) be an \(n \times m\) matrix. Then \[ A = U \Sigma V^T \] where:

  • \(U\) is an \(n \times n\) orthogonal matrix,
  • \(V\) is an \(m \times m\) orthogonal matrix,
  • \(\Sigma\) is an \(n \times m\) rectangular diagonal matrix whose diagonal entries are \[ \sigma_1 \ge \sigma_2 \ge \cdots \ge \sigma_r > 0, \] followed by zeros, where \(\sigma_i\) are the singular values of \(A\).

Properties of SVD

  • Always exists: every matrix (square or rectangular) has an SVD.

  • Singular values come from eigenvalues of \(A^T A\) (or \(A A^T\)), so they are always \(\ge 0\).

  • Connection to eigenvectors:

    • Columns of \(V = [\vec{v}_1, \ldots, \vec{v}_m]\) are orthonormal eigenvectors of \(A^T A\): \[ A^T A \vec{v}_i = \sigma_i^2 \vec{v}_i \]
    • Columns of \(U = [\vec{u}_1, \ldots, \vec{u}_n]\) are orthonormal eigenvectors of \(A A^T\): \[ A A^T \vec{u}_i = \sigma_i^2 \vec{u}_i \]
  • Relationship between \(U\) and \(V\):

    • For each nonzero singular value \(\sigma_i\), \[ \vec{u}_i = \frac{A \vec{v}_i}{\sigma_i}, \qquad \vec{v}_i = \frac{A^T \vec{u}_i}{\sigma_i} \]
    • So \(A\) maps right singular vectors to left singular vectors (up to scaling).
  • \(\Sigma\):

    • tells how much each direction is stretched (by \(\sigma_i\)).
  • Step-by-step geometric picture:

    • \(V^T\): rotate/reflect input (via eigenvectors of \(A^T A\))
    • \(\Sigma\): stretch/shrink
    • \(U\): rotate/reflect output (via eigenvectors of \(A A^T\))
  • Rank connection:

    • number of nonzero singular values = rank of \(A\)
  • Key contrast with diagonalization:

    • diagonalization may fail or not exist
    • SVD always works
    • \(U\) and \(V\) are generally different (unless special cases like symmetric matrices)

Visualizations

Example (3x2 matrix)

Let \[ A=\begin{bmatrix} 1&1\\ 1&0\\ 0&1 \end{bmatrix} \] This is a genuinely nontrivial rectangular matrix: it is neither diagonal nor already in SVD form.

  • Start with \[ A^TA=\begin{bmatrix} 1&1&0\\ 1&0&1 \end{bmatrix} \begin{bmatrix} 1&1\\ 1&0\\ 0&1 \end{bmatrix} = \begin{bmatrix} 2&1\\ 1&2 \end{bmatrix} \]
  • Find eigenvalues of \(A^TA\): \[ \det\begin{bmatrix}2-\lambda&1\\1&2-\lambda\end{bmatrix} =(2-\lambda)^2-1 =\lambda^2-4\lambda+3 =(\lambda-3)(\lambda-1) \]
  • So the eigenvalues are \(\lambda_1=3\) and \(\lambda_2=1\).
  • Therefore the singular values are \[ \sigma_1=\sqrt{3},\qquad \sigma_2=1 \]

Example (3x2 matrix): right singular vectors

  • For \(\lambda_1=3\), solve \[ \begin{bmatrix} -1&1\\ 1&-1 \end{bmatrix} \vec{v}=\vec{0} \] so an eigenvector is \(\begin{bmatrix}1\\1\end{bmatrix}\).
  • Normalize it: \[ \vec{v}_1=\frac{1}{\sqrt2}\begin{bmatrix}1\\1\end{bmatrix} \]
  • For \(\lambda_2=1\), solve \[ \begin{bmatrix} 1&1\\ 1&1 \end{bmatrix} \vec{v}=\vec{0} \] so an eigenvector is \(\begin{bmatrix}1\\-1\end{bmatrix}\).
  • Normalize it: \[ \vec{v}_2=\frac{1}{\sqrt2}\begin{bmatrix}1\\-1\end{bmatrix} \]
  • Thus \[ V=\begin{bmatrix} \frac1{\sqrt2}&\frac1{\sqrt2}\\ \frac1{\sqrt2}&-\frac1{\sqrt2} \end{bmatrix} \]

Example (3x2 matrix): left singular vectors

  • Compute \(\vec{u}_i=\dfrac{A\vec{v}_i}{\sigma_i}\).
  • First, \[ A\vec{v}_1 = \begin{bmatrix} 1&1\\ 1&0\\ 0&1 \end{bmatrix} \frac1{\sqrt2}\begin{bmatrix}1\\1\end{bmatrix} = \frac1{\sqrt2}\begin{bmatrix}2\\1\\1\end{bmatrix} \] so \[ \vec{u}_1=\frac{1}{\sqrt3}\cdot \frac1{\sqrt2}\begin{bmatrix}2\\1\\1\end{bmatrix} =\frac1{\sqrt6}\begin{bmatrix}2\\1\\1\end{bmatrix} \]
  • Next, \[ A\vec{v}_2 = \begin{bmatrix} 1&1\\ 1&0\\ 0&1 \end{bmatrix} \frac1{\sqrt2}\begin{bmatrix}1\\-1\end{bmatrix} = \frac1{\sqrt2}\begin{bmatrix}0\\1\\-1\end{bmatrix} \] so \[ \vec{u}_2=\frac1{\sqrt2}\begin{bmatrix}0\\1\\-1\end{bmatrix} \]
  • To make \(U\) a full \(3\times3\) orthogonal matrix, choose a unit vector orthogonal to both \(\vec{u}_1,\vec{u}_2\): \[ \vec{u}_3=\frac1{\sqrt6}\begin{bmatrix}-2\\1\\1\end{bmatrix} \]

Example (3x2 matrix): final SVD

  • Hence \[ U= \begin{bmatrix} \frac2{\sqrt6}&0&-\frac2{\sqrt6}\\ \frac1{\sqrt6}&\frac1{\sqrt2}&\frac1{\sqrt6}\\ \frac1{\sqrt6}&-\frac1{\sqrt2}&\frac1{\sqrt6} \end{bmatrix}, \qquad \Sigma= \begin{bmatrix} \sqrt3&0\\ 0&1\\ 0&0 \end{bmatrix}, \qquad V= \begin{bmatrix} \frac1{\sqrt2}&\frac1{\sqrt2}\\ \frac1{\sqrt2}&-\frac1{\sqrt2} \end{bmatrix} \]
  • Therefore \[ A=U\Sigma V^T \]
  • Geometrically:
    • \(V^T\) rotates the input coordinates in \(\mathbb R^2\)
    • \(\Sigma\) stretches by \(\sqrt3\) in one direction and by \(1\) in the perpendicular direction
    • \(U\) rotates the result into \(\mathbb R^3\)

Example (nondiagonalizable 2x2 matrix)

Let \[ A=\begin{bmatrix} 1&1\\ 0&1 \end{bmatrix} \] This matrix has only one eigenvector, so it is not diagonalizable.

  • To see that quickly, note that the only eigenvalue is \(1\) since \[ \det(A-\lambda I)=\det\begin{bmatrix}1-\lambda&1\\0&1-\lambda\end{bmatrix}=(1-\lambda)^2 \]
  • But \[ A-I=\begin{bmatrix}0&1\\0&0\end{bmatrix} \] has one-dimensional nullspace, so there is only one eigenvector direction.
  • Even though diagonalization fails, SVD still exists.
  • Compute \[ A^TA= \begin{bmatrix} 1&0\\ 1&1 \end{bmatrix} \begin{bmatrix} 1&1\\ 0&1 \end{bmatrix} = \begin{bmatrix} 1&1\\ 1&2 \end{bmatrix} \]

Example (nondiagonalizable 2x2 matrix): singular values

  • Find the eigenvalues of \(A^TA\): \[ \det\begin{bmatrix}1-\lambda&1\\1&2-\lambda\end{bmatrix} =(1-\lambda)(2-\lambda)-1 =\lambda^2-3\lambda+1 \]
  • Thus \[ \lambda_{1,2}=\frac{3\pm\sqrt5}{2} \]
  • Let \[ \phi=\frac{1+\sqrt5}{2} \] Then \[ \lambda_1=\phi^2,\qquad \lambda_2=\phi^{-2} \]
  • Therefore the singular values are \[ \sigma_1=\phi,\qquad \sigma_2=\phi^{-1} \]

Example (nondiagonalizable 2x2 matrix): singular vectors

  • For \(\lambda_1=\phi^2\), solve \[ (1-\lambda_1)x+y=0 \] so we may take an eigenvector \(\begin{bmatrix}1\\\phi\end{bmatrix}\).
  • For \(\lambda_2=\phi^{-2}\), a perpendicular eigenvector is \(\begin{bmatrix}\phi\\-1\end{bmatrix}\).
  • After normalization, \[ \vec{v}_1=\frac1{\sqrt{1+\phi^2}}\begin{bmatrix}1\\\phi\end{bmatrix}, \qquad \vec{v}_2=\frac1{\sqrt{1+\phi^2}}\begin{bmatrix}\phi\\-1\end{bmatrix} \]
  • Now compute \(\vec{u}_i=\dfrac{A\vec{v}_i}{\sigma_i}\): \[ A\vec{v}_1 = \frac1{\sqrt{1+\phi^2}} \begin{bmatrix} 1+\phi\\ \phi \end{bmatrix} = \frac{\phi}{\sqrt{1+\phi^2}} \begin{bmatrix} \phi\\ 1 \end{bmatrix} \] so \[ \vec{u}_1=\frac1{\sqrt{1+\phi^2}} \begin{bmatrix} \phi\\ 1 \end{bmatrix} \]

Example (nondiagonalizable 2x2 matrix): final SVD

  • Similarly, \[ A\vec{v}_2 = \frac1{\sqrt{1+\phi^2}} \begin{bmatrix} \phi-1\\ -1 \end{bmatrix} = \frac1{\sqrt{1+\phi^2}} \begin{bmatrix} \phi^{-1}\\ -1 \end{bmatrix} \] so \[ \vec{u}_2=\frac1{\sqrt{1+\phi^2}} \begin{bmatrix} 1\\ -\phi \end{bmatrix} \]
  • Therefore \[ U= \frac1{\sqrt{1+\phi^2}} \begin{bmatrix} \phi&1\\ 1&-\phi \end{bmatrix}, \qquad \Sigma= \begin{bmatrix} \phi&0\\ 0&\phi^{-1} \end{bmatrix}, \qquad V= \frac1{\sqrt{1+\phi^2}} \begin{bmatrix} 1&\phi\\ \phi&-1 \end{bmatrix} \]
  • So \[ A=U\Sigma V^T \]
  • This is a good comparison point:
    • diagonalization fails
    • SVD still works perfectly

Example (nondiagonalizable 3x3 matrix)

Let \[ A=\begin{bmatrix} 0&1&1\\ 0&0&1\\ 0&0&0 \end{bmatrix} \] This is not diagonalizable because it is nonzero nilpotent: \(A^3=0\), and a nonzero nilpotent matrix cannot have a basis of eigenvectors.

  • Compute \[ A^TA= \begin{bmatrix} 0&0&0\\ 1&0&0\\ 1&1&0 \end{bmatrix} \begin{bmatrix} 0&1&1\\ 0&0&1\\ 0&0&0 \end{bmatrix} = \begin{bmatrix} 0&0&0\\ 0&1&1\\ 0&1&2 \end{bmatrix} \]
  • The lower-right \(2\times2\) block is exactly the same matrix as in the previous example.
  • So the eigenvalues are \[ 0,\ \phi^2,\ \phi^{-2} \]
  • Hence the singular values are \[ \sigma_1=\phi,\qquad \sigma_2=\phi^{-1},\qquad \sigma_3=0 \]

Example (nondiagonalizable 3x3 matrix): singular vectors

  • Right singular vectors can be chosen as \[ \vec{v}_1=\frac1{\sqrt{1+\phi^2}} \begin{bmatrix} 0\\ 1\\ \phi \end{bmatrix}, \qquad \vec{v}_2=\frac1{\sqrt{1+\phi^2}} \begin{bmatrix} 0\\ \phi\\ -1 \end{bmatrix}, \qquad \vec{v}_3=\begin{bmatrix}1\\0\\0\end{bmatrix} \]
  • For the nonzero singular values, \[ \vec{u}_i=\frac{A\vec{v}_i}{\sigma_i} \]
  • Compute \[ A\vec{v}_1 = \frac1{\sqrt{1+\phi^2}} \begin{bmatrix} 1+\phi\\ \phi\\ 0 \end{bmatrix} = \frac{\phi}{\sqrt{1+\phi^2}} \begin{bmatrix} \phi\\ 1\\ 0 \end{bmatrix} \] so \[ \vec{u}_1=\frac1{\sqrt{1+\phi^2}} \begin{bmatrix} \phi\\ 1\\ 0 \end{bmatrix} \]

Example (nondiagonalizable 3x3 matrix): final SVD

  • Likewise, \[ A\vec{v}_2 = \frac1{\sqrt{1+\phi^2}} \begin{bmatrix} \phi-1\\ -1\\ 0 \end{bmatrix} = \frac1{\sqrt{1+\phi^2}} \begin{bmatrix} \phi^{-1}\\ -1\\ 0 \end{bmatrix} \] so \[ \vec{u}_2=\frac1{\sqrt{1+\phi^2}} \begin{bmatrix} 1\\ -\phi\\ 0 \end{bmatrix} \]
  • Since rank\((A)=2\), choose the last left singular vector as \[ \vec{u}_3=\begin{bmatrix}0\\0\\1\end{bmatrix} \]
  • Thus \[ U= \begin{bmatrix} \frac{\phi}{\sqrt{1+\phi^2}}&\frac{1}{\sqrt{1+\phi^2}}&0\\ \frac{1}{\sqrt{1+\phi^2}}&-\frac{\phi}{\sqrt{1+\phi^2}}&0\\ 0&0&1 \end{bmatrix}, \qquad \Sigma= \begin{bmatrix} \phi&0&0\\ 0&\phi^{-1}&0\\ 0&0&0 \end{bmatrix} \]
  • Again, SVD exists and is useful even when diagonalization fails badly.

Example (diagonalizable 2x2 matrix)

Let \[ A=\begin{bmatrix} 2&1\\ 1&2 \end{bmatrix} \] This one is symmetric, so it is diagonalizable by an orthogonal matrix.

  • Compute the characteristic polynomial: \[ \det\begin{bmatrix} 2-\lambda&1\\ 1&2-\lambda \end{bmatrix} =(2-\lambda)^2-1 =(\lambda-3)(\lambda-1) \]
  • So the eigenvalues are \(3\) and \(1\).
  • Corresponding orthonormal eigenvectors are \[ \vec{u}_1=\frac1{\sqrt2}\begin{bmatrix}1\\1\end{bmatrix}, \qquad \vec{u}_2=\frac1{\sqrt2}\begin{bmatrix}1\\-1\end{bmatrix} \]
  • Therefore the diagonalization is \[ A=PDP^T \] with \[ P=\begin{bmatrix} \frac1{\sqrt2}&\frac1{\sqrt2}\\ \frac1{\sqrt2}&-\frac1{\sqrt2} \end{bmatrix}, \qquad D=\begin{bmatrix} 3&0\\ 0&1 \end{bmatrix} \]

Example (diagonalizable 2x2 matrix): compare with SVD

  • Because \(A\) is symmetric and its eigenvalues are all positive,
    • the singular values are exactly the eigenvalues,
    • the left and right singular vectors are the same.
  • So the SVD is \[ A=U\Sigma V^T \] with \[ U=V=P, \qquad \Sigma=D= \begin{bmatrix} 3&0\\ 0&1 \end{bmatrix} \]
  • In this special case, diagonalization and SVD look essentially the same.
  • This is the exception, not the rule.

Example (diagonalizable 3x3 matrix)

Let \[ A=\begin{bmatrix} 2&1&0\\ 0&1&0\\ 0&0&3 \end{bmatrix} \] This matrix is diagonalizable because it has three distinct eigenvalues \(2,1,3\).

  • Since the diagonal entries are distinct, the eigenvalues are immediately \[ 2,\ 1,\ 3 \]
  • Eigenvectors:
    • for \(\lambda=2\), \(\vec{p}_1=\begin{bmatrix}1\\0\\0\end{bmatrix}\)
    • for \(\lambda=1\), solve \((A-I)\vec{x}=\vec{0}\), giving \(\vec{p}_2=\begin{bmatrix}-1\\1\\0\end{bmatrix}\)
    • for \(\lambda=3\), \(\vec{p}_3=\begin{bmatrix}0\\0\\1\end{bmatrix}\)
  • So \(A=PDP^{-1}\) with \[ D=\mathrm{diag}(2,1,3) \]
  • Now compare this with SVD.

Example (diagonalizable 3x3 matrix): singular values

  • Compute \[ A^TA= \begin{bmatrix} 2&0&0\\ 1&1&0\\ 0&0&3 \end{bmatrix} \begin{bmatrix} 2&1&0\\ 0&1&0\\ 0&0&3 \end{bmatrix} = \begin{bmatrix} 4&2&0\\ 2&2&0\\ 0&0&9 \end{bmatrix} \]
  • The lower-right entry already gives one eigenvalue \(9\), hence one singular value \(3\).
  • For the upper-left \(2\times2\) block, \[ \det\begin{bmatrix} 4-\lambda&2\\ 2&2-\lambda \end{bmatrix} =(4-\lambda)(2-\lambda)-4 =\lambda^2-6\lambda+4 \]
  • Thus the other two eigenvalues are \[ \lambda=3\pm\sqrt5 \]
  • Therefore the singular values are \[ \sigma_1=3,\qquad \sigma_2=\sqrt{3+\sqrt5},\qquad \sigma_3=\sqrt{3-\sqrt5} \]

Example (diagonalizable 3x3 matrix): comparison

  • Notice the key point:
    • eigenvalues of \(A\) are \(2,1,3\)
    • singular values of \(A\) are \(3,\sqrt{3+\sqrt5},\sqrt{3-\sqrt5}\)
  • These are not the same.
  • So even though \(A\) is diagonalizable, its diagonalization and SVD are describing different things:
    • diagonalization uses eigenvectors of \(A\)
    • SVD uses orthonormal singular vectors coming from \(A^TA\) and \(AA^T\)
  • Moral:
    • for symmetric positive matrices, diagonalization and SVD may agree
    • for general matrices, even diagonalizable ones, SVD is a different decomposition